Skip to content

[Refactor][TCPCG] Simplify quantization paths after TCPCG retirement (9/9) - #42293

Open
Oasis-Git wants to merge 10 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-09-quantization
Open

Oasis-Git wants to merge 10 commits into
sgl-project:mainfrom
Oasis-Git:refactor/tcpcg-09-quantization

Conversation

@Oasis-Git

@Oasis-Git Oasis-Git commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This is step 9/9, split from the original TCPCG removal PR #41634.

Depends on #42292 (8/9); merge those dependencies first. This branch targets upstream main, so its diff includes unmerged prerequisites. Review this step alone using the isolated step-9 diff.

Series

  1. Relocate shared graph tensor and DSA head-gate helpers — [Refactor][TCPCG] Relocate shared graph tensor and DSA head-gate helpers (1/9) #42285.
  2. Add explicit batch support to BCG eager regions — [Refactor][TCPCG] Bind live forward batches in BCG eager calls (2/9) #42286.
  3. Retire TCPCG backend selection, configuration, documentation, and tests — [Refactor][TCPCG] Retire backend selection and configuration (3/9) #42287.
  4. Consolidate Radix attention eager regions and isolate Inkling-specific behavior — [Refactor][TCPCG] Consolidate attention eager regions and isolate Inkling policy (4/9) #42288.
  5. Migrate DSA and DeepSeek eager regions — [Refactor][TCPCG] Migrate DSA and DeepSeek graph breaks to eager methods (5/9) #42289.
  6. Migrate MoE and Mamba eager regions — [Refactor][TCPCG] Migrate MoE and Mamba graph execution to eager methods (6/9) #42290.
  7. Migrate diffusion eager wrappers — [Refactor][TCPCG] Standardize diffusion eager graph wrappers (7/9) #42291. Depends only on step 2.
  8. Remove obsolete runtime contexts, registries, and split-op plumbing — [Refactor][TCPCG] Remove obsolete runtime registries and split-op plumbing (8/9) #42292.
  9. Simplify quantization paths separately because of their torch.compile implications — [Refactor][TCPCG] Simplify quantization paths after TCPCG retirement (9/9) #42293.

The disputed relocation of inactive compiler/reference files is excluded pending a separate decision.

Changes

Remove TCPCG-specific quantization indirection after the runtime migration.

  • Remove Marlin's legacy context lookup and associated custom-op adapters.
  • Simplify FP8 helpers, gated normalization, and related model utility paths.
  • Update the FP8 helper test for the resulting interface.

Keep this change separate because removing opaque custom-op boundaries can affect torch.compile behavior. This PR does not remove the torch.compile framework.

The cumulative series matches the original refactor except for the deliberately excluded reference-folder relocation. Inactive TCPCG files remain at their original paths; the final-tree scan found no imports of them from active Python code.

Validation

Pre-commit and Python syntax checks passed. The final combined CPU suite passed 68 tests plus 16 subtests using the macOS import shim. Those tests do not validate quantization kernels or their torch.compile behavior: GPU quantization, compilation, model accuracy, and upstream CI remain to be validated.


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 deepseek diffusion SGLang Diffusion documentation Improvements or additions to documentation jit-kernel Multi-modal multi-modal language model npu piecewise-cuda-graph speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant